Papers with general reasoning and mathematical benchmarks

    1 papers
    ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection (2026.findings-acl)

    Copied to clipboard

    Challenge: Traditional fine-tuning ignores one-to-many nature of language, leading to overfitting . authors propose a method to fine- tune LLMs by leveraging tokens.
    Approach: They propose a method to fine-tune Large Language Models by leveraging tokens to mask low-probability tokens.
    Outcome: The proposed method outperforms baselines on general reasoning and mathematical benchmarks.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations